[NVBUG-6448152][test] TEST ONLY pre-cancellation native-source discriminator#16854
[NVBUG-6448152][test] TEST ONLY pre-cancellation native-source discriminator#16854chienchunhung wants to merge 5 commits into
Conversation
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
…sensus factor Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
Signed-off-by: Chien-Chun Hung <2679986+chienchunhung@users.noreply.github.com>
|
/bot run --disable-fail-fast --stage-list "GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1" |
|
PR_Github #61657 [ run ] triggered by Bot. Commit: |
|
PR_Github #61657 [ run ] completed with state
|
|
Terminal evidence for this TEST ONLY discriminator:
This is +0.40% versus the later-tree 794.61 result and 51.51% of the historical 1548.84 result. There is no throughput recovery at the native 544199a4 checkpoint, so the following three first-parent commits through 58d8964 are excluded and the slow bound moves earlier. The diagnostic branch is preserved for reproducibility; this TEST ONLY draft is now complete and will be closed. |
This is one fixed-current-runtime source discriminator for the separate C++ CTX PP forward/device-loop throughput regression. The asynchronous-consensus product change is already merged and is not under review here.
One-factor construction
544199a47b91b65d295322c6bb1f74dadbcf6a795726e87af6bc2267e116eb26c1bd71ef9c5fded6f665e59a7d650672433743bb6233712e04c5a90f58d8964dcheckpoint from the preceding discriminator. The two intervening commits before58d8964dare test-waiver and CODEOWNERS changes;58d8964dis the only intervening product-code change.202607151440-16194; stable patch ID:7e73140158499de670b40490317c6f77601d06fc.52adfb97177ae01283e352d36d52b428f2f1bf00.mainis present only as a no-tree second parent for CI mergeability.Interpretation
58d8964d, where the only product-code change is the NIXL in-flight-cancellation change.Requested evidence
Run only
GB300-12_GPUs-3_Nodes-PyTorch-Disagg-PerfSanity-CTX1-NODE1-GPU4-GEN1-NODE2-GPU8-Post-Merge-1, requiring the exact selector, three nodes, twelve tasks, four GPUs per node,--no-container-mount-home, 512/512 successful requests, coordinator activation with protocol mode 2 and cancellation disabled on all four CTX PP ranks, no timeout/cancellation events, clean shutdown, and a valid official throughput metric.The draft will be closed after terminal evidence is captured; the branch will be preserved.